Papers with document analysis
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering (2024.naacl-demo)
Copied to clipboard
| Challenge: | DOCMASTER is a platform for annotating PDF documents, model training, and inference, tailored to document question-answering. |
| Approach: | They propose to integrate layout information into a unified platform for annotating PDF documents, model training, and inference tailored to document question-answering. |
| Outcome: | The proposed platform is designed for annotating PDF documents, model training, and inference, tailored to document question-answering. |
IFlyLegal: A Chinese Legal System for Consultation, Law Searching, and Document Analysis (D19-3)
Copied to clipboard
| Challenge: | Legal Tech is a system that performs legal consulting, multi-way law searching, and legal document analysis using deep contextual representations and various attention mechanisms. |
| Approach: | They propose a Chinese legal system that performs legal consulting, multi-way law searching, and legal document analysis using deep contextual representations and various attention mechanisms. |
| Outcome: | The proposed system performs legal consulting, multi-way law searching, and legal document analysis by exploiting techniques such as deep contextual representations and various attention mechanisms. |
Lessons from the Bible on Modern Topics: Low-Resource Multilingual Topic Model Evaluation (N18-1)
Copied to clipboard
| Challenge: | Existing metrics to evaluate multilingual topic quality are inadequate for multilingual document analysis. |
| Approach: | They propose a new intrinsic evaluation metric for multilingual topic models that correlates well with human judgments of multilingual coherence and performance in downstream applications. |
| Outcome: | The proposed model improves the performance of multilingual topic models in low-resource languages and with human judgments of multilinguistic topic coherence. |
Large Language Models Are Still Misled by Simple Bias Ensembles (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing benchmarks for large language models are constrained to datasets where each sample is manually injected with only one type of bias. |
| Approach: | They propose a multi-bias benchmark where each sample contains multiple types of biases. |
| Outcome: | The proposed benchmark shows that existing LLMs and debiasing methods perform poorly on this benchmark, highlighting the challenge of eliminating compounded biases. |
Squeezed Attention: Accelerating Long Context Length LLM Inference (2025.acl-long)
Copied to clipboard
Coleman Richard Charles Hooper, Sehoon Kim, Hiva Mohammadzadeh, Monishwaran Maheswaran, Sebastian Zhao, June Paik, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
| Challenge: | Emerging Large Language Models require long input context to perform complex tasks. |
| Approach: | They propose an algorithm to reduce the complexity of attention with respect to the fixed context length. |
| Outcome: | The proposed method reduces the complexity of attention from linear to logarithmic with respect to the fixed context length. |
Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing work exhibits biases toward English and prepositional marking . Existing models are limited in understanding spatial relations across typologically diverse languages . |
| Approach: | They propose a multilingual framework and benchmark for spatial language understanding . they decompose spatial relations into surface elements and semantic components . their results suggest surface parsing does not entail spatial understanding - they argue . |
| Outcome: | The proposed framework and benchmark decomposes spatial relations into surface elements and semantic components. |